Papers with Evaluation metrics
MapQaTor: An Extensible Framework for Efficient Annotation of Map-Based QA Datasets (2025.acl-demo)
Copied to clipboard
| Challenge: | Mapping and navigation services struggle to handle natural language geospatial queries. |
| Approach: | They introduce an extensible open-source framework that streamlines the creation of reproducible, traceable map-based QA datasets. |
| Outcome: | a new open-source framework streamlines the creation of reproducible, traceable map-based QA datasets. |
SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods for Open Domain Event Detection (ODED) lack representative representations of the real world, making it difficult to accurately reflect performance of various ODED methods in real-world scenarios. |
| Approach: | They propose a scalable and reliable Semantic-level Evaluation framework for Open domain event detection by constructing a more representative evaluation benchmark and introducing a semantic evaluation metric. |
| Outcome: | The proposed framework first constructs a more representative evaluation benchmark that currently includes 564 event types covering 7 major domains, with a cost-effective supplementary annotation strategy to ensure the benchmark’s representativeness. |
Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors (2021.emnlp-main)
Copied to clipboard
| Challenge: | Evaluation metrics are a key ingredient for progress of text generation systems . a class of novel evaluation metrics based on BERT and its variants has been explored . |
| Approach: | They propose to disentangle BERT-based evaluation metrics along linguistic factors . they show they are sensitive to lexical overlap, just like BLEU and ROUGE . |
| Outcome: | The proposed metrics capture all aspects but are sensitive to lexical overlap, just like BLEU and ROUGE, the authors show . |
IM^2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Evaluation metrics for dialogue systems are expensive and time-consuming . current evaluation metrics focus on a single quality or several qualities . |
| Approach: | They propose an interpretable, multi-faceted, and controllable framework to combine dialogue metrics which are good at measuring different qualities. |
| Outcome: | The proposed framework integrates a large number of evaluation metrics to improve the performance of the model. |